Papers with RL methods

15 papers
Learning Natural Language Generation with Truncated Reinforcement Learning (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to train conditional languagemodels without supervised learning fail to scale to large action spaces, thus allowing to train a language agent by only interacting with its environment without any task-specific prior knowledge.
Approach: They propose an original approach to train conditional languagemodels without supervised learning by only using reinforcement learning.
Outcome: The proposed approach avoids the dependency to labelled datasets and reduces pretrained policy flaws such as language or exposure biases.
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings (2026.tacl-1)

Copied to clipboard

Challenge: Reinforcement learning (RL) is an effective and robust method for training neural machine translation systems.
Approach: They propose a method that leverages fine-grained, token-level quality assessments . they use a state-of-the-art quality estimation system as their token- level reward model .
Outcome: The proposed approach leverages fine-grained, token-level quality assessments along with error severity levels to improve translation quality.
Reinforcement Learning with Token-level Feedback for Controllable Text Generation (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for controllable text generation are guided by coarse-grained feedback, which may lead to suboptimal performance owing to semantic twists or progressions within sentences.
Approach: They propose a reinforcement learning algorithm which formulates TOken-LEvel rewards for controllable text generation and employs a "first-quantize-then-noise" paradigm to enhance the robustness of the RL algorithm.
Outcome: The proposed algorithm can achieve superior performance on single-attribute and multi-attract control tasks.
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in diffusion models have shown impressive performance in many domains, but their ability to follow instructions is still unsatisfactory.
Approach: They propose an algorithm that aligns images to text through iterative image sampling and prompt relabeling with feedback.
Outcome: The proposed algorithm improves on the spatial relation VISOR benchmark by 15.22% compared to previous methods.
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing RL methods rely on unstructured self-sampling to fit scalar rewards, resulting in inefficient rollouts.
Approach: They propose a structured template-guided RL framework that augments policy optimization with explicit template guidance.
Outcome: Experiments show that TemplateRL outperforms GRPO and GRPI by 99% on AIME and 41% on AMC with superior stability on weak models and remarkable cross-domain generalization.
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization (2025.emnlp-main)

Copied to clipboard

Challenge: Recent research focuses on integrating reasoning capabilities into the realm of retrieval-augmented generation (RAG) via outcome-supervised reinforcement learning (RL).
Approach: They propose a process-level reward module to mitigate the unawareness of intermediate reasoning steps in outcome-level supervision without additional annotation.
Outcome: The proposed framework can boost LLMs’ reasoning ability by integrating external knowledge sources through retrieval-augmented generation (RAG) The proposed model can mitigate the unawareness of intermediate reasoning steps in outcome-level supervision without additional annotation.
Distill and Align Decomposition for Enhanced Claim Verification (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for complex claim verification struggle to align decomposition quality with verification performance.
Approach: They propose a reinforcement learning approach that optimizes decomposition quality and verifier alignment using Group Relative Policy Optimization.
Outcome: The proposed method outperforms prompt-based approaches and existing methods in six evaluation settings.
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint (2024.findings-acl)

Copied to clipboard

Challenge: Existing reinforcement learning methods do not provide fine-grained supervision for complex reasoning tasks.
Approach: They propose a reinforcement learning method that incorporates a generative model as the reward model and a token-level supervision model for RL training.
Outcome: Experiments on 8 tasks show the proposed method is effective .
A Teacher-Student Framework for Maintainable Dialog Manager (D18-1)

Copied to clipboard

Challenge: Reinforcement learning (RL) is an attractive solution for task-oriented dialog systems . but extending RL-based systems to handle new intents and slots requires a system redesign .
Approach: They propose a teacher-student framework to extend RL-based dialog systems . they propose to specify constraints held in the new dialog manager .
Outcome: The proposed framework makes no assumption about unsupported intents and slots, making it possible to improve RL-based systems incrementally.
LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated emergent common-sense reasoning and Theory of Mind (ToM) capabilities, making them promising candidates for developing coordination agents.
Approach: They propose to use Large Language Models (LLMs) to analyze coordination models in Pure Coordination settings where agents must cooperate to maximize gains.
Outcome: The proposed benchmark evaluates LLMs through two distinct tasks: Agentic Coordination and Coordination Question Answering.
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing RL methods focus on generation tasks while neglecting dialogue state tracking (DST) for understanding.
Approach: They propose a method that integrates RL into both understanding and generation tasks by introducing step-by-step rewards throughout the token generation.
Outcome: The proposed approach achieves state-of-the-art results on three widely used datasets.
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data (2024.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement Learning from Human Feedback (RLHF) relies on complex methodologies like Proximal Policy Optimization (PPO) that require extensive hyper-parameter tuning and pose challenges in sample efficiency and stability.
Approach: They propose an innovative framework that leverages direct preference optimization techniques but extends them by estimating the conditionally optimal policy directly from the model’s responses.
Outcome: The proposed framework matches and exceeds the effectiveness of Proximal Policy Optimization (PPO) in terms of convergence speed and alignment of model responses with human preferences.
Efficient (Soft) Q-Learning for Text Generation with Limited Good Data (2022.findings-emnlp)

Copied to clipboard

Challenge: Maximum likelihood estimation (MLE) is the predominant method for training text generation models.
Approach: They propose a new RL formulation for text generation from the soft Q-learning perspective using path consistency learning to combine the best of on-/off-policy updates and learn effectively from sparse reward.
Outcome: The proposed approach outperforms MLE and previous RL methods in a wide range of tasks.
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for reinforcement learning for large language models do not accurately assess generalization.
Approach: They propose three core principles for designing more faithful benchmarks: sufficient difficulty, balanced evaluation, and distributional robustness.
Outcome: The proposed benchmarks do not accurately assess generalization across distribution shifts, difficulty levels, and counterfactual scenarios.
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization (2026.acl-long)

Copied to clipboard

Challenge: Current reinforcement learning methods suffer from coarse-grained, trajectory-level rewards that provide insufficient learning signals for complex multi-turn interactions, leading to training stagnation.
Approach: They propose a novel RL algorithm for training large language models for multi-turn tool-integrated reasoning (TIR) that incorporates three innovations: turn-level reward assignment that provides fine-grained feedback for individual turns, return-based advantage estimation where normalized discounted returns are calculated as advantages, and self-supervised reward shaping that exploits self-supervision signals from generated code to densify sparse binary outcome-based rewards.
Outcome: The proposed algorithm outperforms GRPO by 3.0% across diverse math reasoning benchmarks and improves grepo by 3.9% on commonsense reasoning and program synthesis tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations